Papers by Marta Gonzalez Mallo

1 papers
Automatic Evaluation of Healthcare LLMs Beyond Question-Answering (2025.naacl-short)

Copied to clipboard

Challenge: Current Large Language Models (LLMs) benchmarks are often based on open-ended or close-ended QA evaluations, avoiding the requirement of human labor.
Approach: They propose a multi-axis suite for healthcare LLM evaluation, exploring correlations between open and close benchmarks and metrics.
Outcome: The proposed framework explores correlations between open and close benchmarks and metrics in the healthcare domain, with blind spots and overlaps in existing methodologies.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations